Papers with semantic similarity scores

5 papers
Exploring the Semantic Space of Second Language Learners (2026.eacl-srw)

Copied to clipboard

Challenge: Using machine learning models, we compared the semantic space of university-level students learning French with native speakers' (L1) .
Approach: They extracted semantic features from narrative text and used interpretability techniques to identify the most informative features per model.
Outcome: The results show that the second language learners had higher semantic similarity scores than the native speakers at the token level, whereas the similarity decreased over time but did not reach native-level values.
Sentence Similarity Based on Contexts (2022.tacl-1)

Copied to clipboard

Challenge: Existing methods to measure sentence similarity face limited dataset size and training-test gap . existing methods lack large-scale labeled datasets with labeles that are labor-intensive and expensive .
Approach: They propose a framework that measures sentence similarity by comparing probabilities of generating two sentences given the same context.
Outcome: The proposed framework achieves significant performance boosts over baselines under supervised and unsupervised settings.
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: generative Large Language Models (LLMs) are based on natural text evolution .
Approach: They propose a framework for curating naturally evolved variants of reading passages from contemporary QA benchmarks and for analysing LLM performance across a range of semantic similarity scores.
Outcome: The proposed framework evaluates QA datasets and LLMs with publicly available training data.
Distractor Generation Using Generative and Discriminative Capabilities of Transformer-based Models (2024.lrec-main)

Copied to clipboard

Challenge: Multiple Choice Questions (MCQs) are used to test language learners' comprehension and knowledge.
Approach: They propose an automatic distractor generation approach which generates correct and incorrect answer options and then discriminates potential correct options from distractors.
Outcome: The proposed approach outperforms previous models on multiple choice questions and reading comprehension questions.
Mapping semantic networks to Dutch word embeddings as a diagnostic tool for cognitive decline (2025.emnlp-main)

Copied to clipboard

Challenge: Semantic networks are abstract representations of the semantic memory system and can be used to estimate networks .
Approach: They used Dutch verbal fluency data to explore the relationship between semantic networks and cognitive health.
Outcome: The proposed measures predict cognitive health scores on the Mini-Mental State Examination (MMSE) while the traditional number-of-words measure was not significant, the results suggest that semantic network metrics may provide a more sensitive measure of cognitive health than traditional scoring.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations